New Issue: Orbital Catastrophe Ahead? Read Now

OpenAI’s latest math breakthroughs commit research misconduct, experts say

OpenAI released 10 AI-generated results over the weekend. Some mathematicians are unhappy with their approach

An illustration of a robot copying a human's abstract painting.
Malte Mueller/Getty Images

OpenAI’s newest chatbot may be a whiz at math, but it seems to be lagging far behind humans in its academic rigor.

Last week the company announced 10 more artificial-intelligence-generated math advances that were found during internal development and testing of its next major large language model. This batch of results came from that LLM, Astra, and each one resolves or progresses a different “long-standing open problem” of “substantial interest” to the mathematical community. The company said that the total token cost was a mere $2,000.

The news quickly spread as yet another harbinger of AI’s promise—or threat—of outpacing humans to become a dominant, disruptive force in math and computer science. But over the days following the announcement, as flesh-and-blood experts pored over the nearly 250-page paper in detail, many grew frustrated.


On supporting science journalism

If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.


Two of the most exciting results, the experts say, incorporate preexisting ideas from the recent mathematical literature without properly citing them. This contradicts OpenAI’s initial press release, which said that the problems Astra addressed “have been open and seen no progress on the main result for at least a decade.” (OpenAI has since updated the language to be more accurate).

“They are running roughshod over the work of others who came before them in a deliberate way,” says Stephen Miller, a mathematician at Yeshiva University, who argues that OpenAI has effectively plagiarized his own research. “It seems completely systematic to me, and it points to research misconduct.”

The result Miller refers to concerns how many balls you can fit in a box—a seemingly simple problem, except these balls and boxes exist in a mathematical space of 1,000 dimensions—or even more. OpenAI’s paper improves the best estimate for how tightly these balls can possibly be packed. The LLM-generated proof hinges on a particular mathematical argument that it presented as its own but that actually first appeared in a 2016 paper by Miller and a collaborator.

Another of the 10 results resolves a long-standing question in group theory, which studies sets of mathematical objects called “groups” that interact in an organized way. Mathematicians have long wondered whether all groups have a property called “soficity,” the capacity to be faithfully approximated in a particularly way by other, simpler groups. The OpenAI paper establishes at least one group that lacks this property.

The discovery stunned Francesco Fournier-Facio, a mathematician at the University of Cambridge, who studies group theory—at least until he “engaged with this breakthrough as I would if a human had written it,” he says. The result, he and some of his colleagues found, wasn’t as novel as it first appeared. Like a number of recent AI breakthroughs, it pasted together ideas from the mathematical literature to build a new theorem. Once again, the LLM’s trick is its superhuman patience for assembling puzzle pieces, not the ability to make some profound intellectual leap.

In particular, Astra’s key mathematical step combined ideas first found in two papers from 2016 and 2019. Andreas Thom, a mathematician at the Dresden University of Technology, who co-authored the 2019 paper, summarized the result on MathOverflow.com, calling it “creative and at the same time elementary.”

OpenAI’s initial press release seemed to ignore—or be completely unaware of—these crucial, recent developments. Fournier-Facio argues that the two preceding papers show humans had not hit a stalemate with the soficity problem. OpenAI’s mathematicians did their best to attribute these ideas correctly in their paper, he says. But, in spite of their good intentions, “there is the big PR machine that wants to sound as impressive as possible and does not care about being 100 percent accurate,” he says.

“We take responsibility for the correctness of these results and are meeting the same standards generally expected of human mathematicians,” an OpenAI spokesperson said in a statement to Scientific American. “We plan to make small updates [to the paper] this week, consistent with standard academic practice.”

But as AI continues its campaign to conquer math without any built-in fealty to the field’s academic norms, some in the community are clearly losing patience. “OpenAI is now fully participating in high-level research,” Fournier-Facio says. “So they should be held to the same academic standards that we are.”

Editor’s Note (8/17/26): This article was edited after posting to correct the spelling of Stephen Miller’s first name. The text was previously amended on August 7 to correct the description of Andreas Thom co-authoring a 2019 paper.

Subscribe to Support Independent Journalism

Great science journalism requires human expertise, time, effort and creativity. And it costs money. That’s why I and the journalists here at Scientific American hope you’ll join our community.

When you subscribe, you are supporting staff and freelance journalists who are passionate about telling science stories that are true, important and compelling. Our editors and reporters are often experts in their fields, which means they understand the nuances of big discoveries and can untangle the breakthroughs from the hype. With a subscription, you are also supporting rigorous fact-checking to ensure the words we publish are precise and accurate. And you’re supporting original illustrations, graphics and photos that bring you closer to an advanced laboratory, an ice sheet in Antarctica or a space mission in orbit. You’re helping us craft other types of high-quality journalism as well: Our newsletters are carefully written, edited and curated by staffers you have or will come to know and love. Our Science Quickly podcast is based on original reporting, collaboration with editors and scientists and exacting production.

Subscriptions keep this engine running so we can continue to deliver thoughtful, rigorous and independent science journalism to you. In an era of viral misinformation, this work is crucial. If you value what we do, I hope you’ll consider joining us as a subscriber

Thank you,

Jeanna Bryner, Editor in Chief, Scientific American

Subscribe